Papers with MM-IMDb datasets
See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias (2025.naacl-long)
Copied to clipboard
| Challenge: | Vision-language models often rely on a single modality rather than treating and utilizing them equally, leading to dominance of a specific modality on the overall performance. |
| Approach: | They propose a framework to mitigate dominant modality bias by adjusting the gradient of KL divergence based on each modality's contribution and aligning task directions in a non-conflicting manner. |
| Outcome: | The proposed framework mitigates dominant modality bias on UPMC Food-101, Hateful Memes, and MM-IMDb datasets. |
Dynamic Regularization in UDA for Transformers in Multimodal Classification (2023.acl-long)
Copied to clipboard
| Challenge: | Multimodal machine learning is a cutting-edge field that explores ways to combine information from multiple sources into models. |
| Approach: | They propose a multimodal BERT-ViT model that exploits weaker modality while regularizing the loss function. |
| Outcome: | The proposed model exploits weaker modality while regularizing the loss function. |